Papers by Dinh Viet Sang

6 papers
LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing cross-lingual topic models depend on sparse bilingual resources and often yield incoherent or weakly aligned topics.
Approach: They propose a framework that integrates LLM-guided topic refinement with self-consistency uncertainty quantification to enable black-box, stable, and scalable enhancement of cross-lingual topic models.
Outcome: Experiments on multilingual corpora show that the proposed framework achieves superior topic coherence and alignment while reducing reliance on bilingual dictionaries and expensive LLM calls.
XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignments (2025.findings-emnlp)

Copied to clipboard

Challenge: XTRA aims to uncover shared semantic themes across languages . previous methods have achieved improvements in topic diversity but struggle to ensure high topic coherence and consistent alignment across languages.
Approach: a new framework unifies Bag-of-Words modeling with multilingual embeddings is proposed to address this problem . XTRA introduces two core components: (1) representation alignment and (2) topic alignment to enforce cross-lingual consistency.
Outcome: XTRA outperforms baselines in topic coherence, diversity, and alignment quality on multilingual corpora.
FAID: Fine-grained AI-generated Text Detection using Multi-task Auxiliary and Multi-level Contrastive Learning (2026.eacl-long)

Copied to clipboard

Challenge: Existing binary detection frameworks for human-written, LLM-generated and human-LLM collaborative texts are challenging . a recent study focused on binary detection, i.e., human vs. LLM, or on fine-grained detection limited to English.
Approach: They propose a fine-grained detection framework to classify text into three categories . they use multilingual datasets and a multi-domain, multi-generator dataset .
Outcome: The proposed framework outperforms baselines on unseen domains and new LLMs.
Beyond Coherence: Improving Temporal Consistency and Interpretability in Dynamic Topic Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing topic models capture bag-of-words statistics but lack semantic priors . interpretability remains shallow, relying on noisy top-word lists that obscure thematic clarity.
Approach: They propose a variational framework to capture more faithful temporal trajectories . they propose to use entropy-regularized optimal transport to align entire topic constellations .
Outcome: The proposed framework captures more faithful temporal trajectories and improves interpretability.
Multi-Surrogate-Objective Optimization for Neural Topic Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Neural topic modeling incorporates multiple loss functions but can be difficult to optimize for disparate magnitudes of these losses.
Approach: They propose a gradient-based multi-objective optimization approach that integrates MOO algorithms into the model without the need for hard-parameter sharing.
Outcome: The proposed approach outperforms direct MOO applications on NTMs.
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing cross-tokenizer distillation methods are limited by suboptimal alignment across sequence and vocabulary levels.
Approach: They propose a cross-tokenizer distillation framework that enhances token-wise distillation . they use dual-space entropy-based weighting to achieve precise sequence-level alignment .
Outcome: The proposed framework outperforms state-of-the-art methods in large language models but has high computational and memory costs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations